Papers with LLM-as-a-judge framework

2 papers
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards (2025.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions .
Approach: They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples .
Outcome: The proposed framework improves hallucination evaluations by leveraging human-annotated examples.
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances (2026.eacl-long)

Copied to clipboard

Challenge: Common proxies such as Mean Length of Utterance (MLU), lexical diversity (vocd-D), and readability indices are dominated by length and ignore conversational context, missing aspects of response quality such as reasoning depth, topic maintenance, and discourse planning.
Approach: They propose a framework that classifies the Previous Adult Utterance Type and scores the child’s response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’ s contribution to advancing the discourse).
Outcome: The proposed framework assesses the child's response along two axes: Expansion (contextual elaboration and inferential depth) and Independence (the child’s contribution to advancing the discourse).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations